Multi-Level Attention-Based Categorical Emotion Recognition Using Modulation-Filtered Cochleagram
نویسندگان
چکیده
Speech emotion recognition is a critical component for achieving natural human–robot interaction. The modulation-filtered cochleagram feature based on auditory modulation perception, which contains multi-dimensional spectral–temporal representation. In this study, we propose an framework that utilizes multi-level attention network to extract high-level emotional representations from the cochleagram. Our approach channel-level and spatial-level modules generate saliency maps of channel spatial representations, capturing significant space 3D convolution maps, respectively. Furthermore, employ temporal-level module capture regions concatenated sequence maps. experiments Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset demonstrate significantly improves prediction performance categorical compared other evaluated features. Moreover, our achieves comparable unweighted accuracy 71% in by comparing with several existing approaches. summary, study demonstrates effectiveness speech recognition, proposed provides promising direction future research field.
منابع مشابه
Emotion Modulation of Visual Attention: Categorical and Temporal Characteristics
BACKGROUND Experimental research has shown that emotional stimuli can either enhance or impair attentional performance. However, the relative effects of specific emotional stimuli and the specific time course of these differential effects are unclear. METHODOLOGY/PRINCIPAL FINDINGS In the present study, participants (n = 50) searched for a single target within a rapid serial visual presentati...
متن کاملAnalysis of Multi-Lingual Emotion Recognition Using Auditory Attention Features
In this paper, we build mono-lingual and cross-lingual emotion recognition systems and report performance on English and German databases. The emotion recognition system uses biologically inspired auditory attention features together with a neural network for learning the mapping between features and emotion classes. We first build mono-lingual systems for both Berlin Database of Emotional Spee...
متن کاملSpeech Emotion Recognition Using Amplitude Modulation Parameters
In the community of Human Computer Interface (HCI) researchers have been working for several years in trying to emulate a human communication system, using innovative technologies and methodologies, based on the emotion recognition in facial expressions and speech [1-3]. Speech emotion recognition (SER) [4] is a challenging framework in demanding human machine interaction systems. Standard appr...
متن کاملSpeech Emotion Recognition Using Scalogram Based Deep Structure
Speech Emotion Recognition (SER) is an important part of speech-based Human-Computer Interface (HCI) applications. Previous SER methods rely on the extraction of features and training an appropriate classifier. However, most of those features can be affected by emotionally irrelevant factors such as gender, speaking styles and environment. Here, an SER method has been proposed based on a concat...
متن کاملWord-Level Emotion Recognition Using High-Level Features
In this paper, we investigate the use of high-level features for recognizing human emotions at the word-level in natural conversations with virtual agents. Experiments were carried out on the 2012 Audio/Visual Emotion Challenge (AVEC2012) database, where emotions are defined as vectors in the Arousal-Expectancy-Power-Valence emotional space. Our model using 6 novel disfluency features yields si...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
ژورنال
عنوان ژورنال: Applied sciences
سال: 2023
ISSN: ['2076-3417']
DOI: https://doi.org/10.3390/app13116749